Papers with quantitative metrics
Stakeholder Suite: A Unified AI Framework for Mapping Actors, Topics and Arguments in Public Debates (2026.eacl-demo)
Copied to clipboard
| Challenge: | Existing media intelligence tools rely on descriptive analytics with limited transparency. |
| Approach: | They propose a framework for mapping actors, topics, and arguments within public debates . it combines actor detection, topic modeling, argument extraction and stance classification . the framework is tested on multiple energy infrastructure projects as a case study . |
| Outcome: | The proposed framework delivers fine-grained, source-grounded insights while remaining adaptable to diverse domains. |
UFO: A UI-Focused Agent for Windows OS Interaction (2025.naacl-long)
Copied to clipboard
Chaoyun Zhang, Liqun Li, Shilin He, Xu Zhang, Bo Qiao, Si Qin, Minghua Ma, Yu Kang, Qingwei Lin, Saravan Rajmohan, Dongmei Zhang, Qi Zhang
| Challenge: | UFO is a UI-Fcused agent designed to fulfill user requests tailored to Windows OS applications . it decomposes user requests using divide-and-conquer approach, enabling seamless navigation and addressing sub-tasks across multiple applications. |
| Approach: | They propose a UI-Fcused Windows OS agent that decomposes user requests using a divide-and-conquer approach and incorporates a control interaction module tailored for Windows OS. |
| Outcome: | The proposed agent decomposes user requests using divide-and-conquer approach, enabling seamless navigation and addressing sub-tasks across multiple applications. |
Fighting Offensive Language on Social Media with Unsupervised Text Style Transfer (P18-2)
Copied to clipboard
| Challenge: | Existing methods to tackle the problem of offensive language in social media are based on machine learning. |
| Approach: | They propose a method for training encoder-decoders using non-parallel data . they use a collaborative classifier, attention and the cycle consistency loss . |
| Outcome: | The proposed method outperforms state-of-the-art text style transfer systems on Twitter and Reddit . it produces reliable non-offensive transferred sentences, the authors show . |
A Game-Theoretica Negotiation Framework for Cross-Cultural Consensus (2026.acl-long)
Copied to clipboard
| Challenge: | Large language models exhibit pronounced WEIRD cultural bias, marginalizing diverse viewpoints and posing challenges for reconciling diverse populations with varying cultural backgrounds and value systems. |
| Approach: | They propose a framework for cross-cultural fairness using a Nash Equilibrium . they propose equilibriums that iteratively propose and refine natural-language guidelines . |
| Outcome: | The proposed framework generates higher-quality and more balanced consensus . it finetunes diverse LLM architectures with negotiation data, reducing cultural distances by 95.53%. |
T3M: Text Guided 3D Human Motion Synthesis from Speech (2024.findings-naacl)
Copied to clipboard
| Challenge: | Existing methods for speech-driven 3D motion synthesisrely on speech audio . existing methods are inaccurate and inflexible, leading to inflexibility and inefficient synthesis results. |
| Approach: | They propose a text-guided 3D human motion synthesis method that uses text input to generate motions from human speech. |
| Outcome: | The proposed method outperforms existing methods in quantitative and qualitative evaluations. |
ReContraster: Making Your Posters Stand Out with Regional Contrast (2026.acl-long)
Copied to clipboard
| Challenge: | Effective poster design requires rapidly capturing attention and clearly conveying messages. |
| Approach: | They propose a poster-based model that leverages regional contrast to make posters stand out. |
| Outcome: | The proposed model outperforms state-of-the-art methods in producing striking posters. |
QualEval: Qualitative Evaluation for Model Improvement (2024.naacl-long)
Copied to clipboard
Vishvak Murahari, Ameet Deshpande, Peter Clark, Tanmay Rajpurohit, Ashish Sabharwal, Karthik Narasimhan, Ashwin Kalyan
| Challenge: | Quantitative evaluation metrics are inadequate for large language models due to complexity of tasks and cannot provide actionable diagnostics. |
| Approach: | They propose a quantitative evaluation tool called QualEval that uses automated qualitative evaluation as a vehicle for model improvement. |
| Outcome: | The proposed method improves the performance of the Llama 2 model by 15% compared to baselines. |
uMedSum: A Unified Framework for Clinical Abstractive Summarization (2025.acl-long)
Copied to clipboard
Aishik Nagar, Yutong Liu, Andy T. Liu, Viktor Schlegel, Vijay Prakash Dwivedi, Arun-Kumar Kaliya-Perumal, Guna Pratheep Kalanchiam, Yili Tang, Robby T. Tan
| Challenge: | Clinical abstractive summarization struggles to balance faithfulness and informativeness, sacrificing key information or introducing confabulations. |
| Approach: | They develop a modular hybrid framework that integrates confabulation removal and key information addition into abstractive summarization methods. |
| Outcome: | The proposed framework outperforms state-of-the-art abstractive summarization methods in both quantitative metrics and expert evaluations. |
Adaptive Parameterization for Neural Dialogue Generation (D19-1)
Copied to clipboard
| Challenge: | Existing models of open-domain dialogue generate responses based on sequence-to-sequence paradigms. |
| Approach: | They propose an Adaptive Neural Dialogue generation model which manages various conversations with conversation-specific parameterization. |
| Outcome: | The proposed model performs better on a large-scale conversational dataset. |
Evidence > Intuition: Transferability Estimation for Encoder Selection (2022.emnlp-main)
Copied to clipboard
| Challenge: | Existing studies on LM transferability have focused on a priori tuning of encoders . prior work has examined the different yet related tasks of performance prediction . |
| Approach: | They propose to generate quantitative evidence to predict which LM will perform best on a target task without fine-tuning all candidates. |
| Outcome: | The proposed model outperforms the standard of human practitioner ranking in 94% of the setups. |
Agent-Testing Agent: A Meta-Agent for Automated Testing and Evaluation of Conversational AI Agents (2026.eacl-long)
Copied to clipboard
| Challenge: | Robust, developer-friendly evaluation remains a bottleneck. |
| Approach: | They propose a meta-agent that combines static code analysis, developer interrogation, literature mining, and persona-driven adversarial test generation whose difficulty adapts via judge feedback. |
| Outcome: | The agent-testing agent (ATA) surfaces more diverse and severe failures than expert annotators while matching severity, and finishes in 20–30 minutes versus ten-annotator rounds that took days. |
PragmatiCQA: A Dataset for Pragmatic Question Answering in Conversations (2023.findings-acl)
Copied to clipboard
| Challenge: | Mars? - PragmatiCQA |
| Approach: | Mars? - The Paper . |
| Outcome: | The proposed dataset features 6873 QA pairs that explores pragmatic reasoning in conversations over a diverse set of topics. |
WER We Stand: Benchmarking Urdu ASR Models (2025.coling-main)
Copied to clipboard
| Challenge: | This paper analyzes the performance of three ASR models for low-resource languages like Urdu . low-rural languages like urdu have significant gaps in accuracy and reliability . |
| Approach: | They evaluate the performance of three ASR models: Whisper, MMS, and Seamless-M4T . they present the first conversational speech dataset for benchmarking Urdu ASR systems . |
| Outcome: | The proposed model families outperform Whisper, MMS, and Seamless-M4T on two types of speech datasets. |
Multimodal Differential Network for Visual Question Generation (D18-1)
Copied to clipboard
| Challenge: | Current dialog systems show improvement in visual question answering but this does not translate to improved human-AI dialog. |
| Approach: | They propose to use a Multimodal Differential Network to generate natural questions from images using a multimodal differential network. |
| Outcome: | The proposed approach significantly improves over state-of-the-art benchmarks on the quantitative metrics. |
DiplomacyAgent: Do LLMs Balance Interests and Ethical Principles in International Events? (2025.emnlp-main)
Copied to clipboard
| Challenge: | a new study examines the safety implications of large language models in diplomatic positions . it identifies potential risks and ideological biases that could arise from LLMs . |
| Approach: | They propose an LLM-based multi-agent system for diplomatic position analysis . they propose ethical constraint measures to enhance the safety of LLMs . |
| Outcome: | The proposed system assesses the safety implications of large language models in diplomacy . it reveals that LLMs could exhibit a strong bias towards interests, leading to unsafe decisions . |
The rJokes Dataset: a Large Scale Humor Collection (2020.lrec-1)
Copied to clipboard
| Challenge: | Humor is a complex language phenomenon that depends upon many factors, including topic, date, and recipient. |
| Approach: | They compile a large scale humor dataset from the Reddit r/Jokes subreddit. |
| Outcome: | The proposed dataset provides quantitative metrics for the level of humor in each joke, as determined by subreddit user feedback. |
Intrinsic Subgraph Generation for Interpretable Graph Based Visual Question Answering (2024.lrec-main)
Copied to clipboard
| Challenge: | Visual Question Answering (VQA) is acknowledged as a challenging multi-modal task for Machine Learning (ML). |
| Approach: | They propose an interpretable approach for graph-based Visual Question Answering . their model is designed to intrinsically produce a subgraph during the question-answering process as its explanation . |
| Outcome: | The proposed model outperforms existing explainable methods on a graph-based VQA dataset. |
Mitigating Translationese in Low-resource Languages: The Storyboard Approach (2024.lrec-main)
Copied to clipboard
Garry Kuwanto, Eno-Abasi E. Urua, Priscilla Amondi Amuok, Shamsuddeen Hassan Muhammad, Anuoluwapo Aremu, Verrah Otiende, Loice Emma Nanyanga, Teresiah W. Nyoike, Aniefon D. Akpan, Nsima Ab Udouboh, Idongesit Udeme Archibong, Idara Effiong Moses, Ifeoluwatayo A. Ige, Benjamin Ajibade, Olumide Benjamin Awokoya, Idris Abdulmumin, Saminu Mohammad Aliyu, Ruqayya Nasir Iro, Ibrahim Said Ahmad, Deontae Smith, Praise-EL Michaels, David Ifeoluwa Adelani, Derry Tanti Wijaya, Anietie Andy
| Challenge: | Low-resource languages often face challenges in acquiring high-quality language data due to the reliance on translation-based methods, which introduce the translationese effect. |
| Approach: | They propose a method that uses storyboards to elicit more fluent and natural sentences from native speakers without direct exposure to the source text. |
| Outcome: | The proposed method compared with traditional translation-based methods in terms of accuracy and fluency. |
Towards Better Value Principles for Large Language Model Alignment: A Systematic Evaluation and Enhancement (2025.acl-long)
Copied to clipboard
| Challenge: | Large Language Models (LLMs) show remarkable performance across tasks . alignment with human values is critical for their responsible development. |
| Approach: | They propose a framework that evaluates value principles along three desirable properties . they propose supervised fine-tuning, reinforcement learning-based approaches . |
| Outcome: | The proposed framework improves value principles along the three desirable properties of LLMs. |
EducationQ: Evaluating LLMs’ Teaching Capabilities Through Multi-Agent Dialogue Framework (2025.acl-long)
Copied to clipboard
| Challenge: | Large Language Models (LLMs) are increasingly used as educational tools, yet evaluating their teaching capabilities remains challenging due to the resource-intensive nature of teacher-student interactions. |
| Approach: | They propose a multi-agent dialogue framework that efficiently assesses teaching capabilities through simulated dynamic educational scenarios. |
| Outcome: | The proposed framework outperforms open-source models on 1,498 questions across 13 disciplines and 10 difficulty levels on 1,400 questions. |